Skip to main content

Infrastructure

Infrastructure decisions in health are shaped less by technical preference than by three constraints that do not apply elsewhere: data residency law, availability requirements that are clinical safety requirements, and operational capacity that is often the binding limit.

The most sophisticated architecture is worthless if there is nobody to operate it at 3 a.m.


Hosting models​

ModelSuitsCosts
On-premises (facility)Poor connectivity; data must not leave the buildingEvery facility needs power, hardware, backup and someone to maintain them
On-premises (national data centre)Residency requirements; existing government capacityCapital-intensive; scaling is procurement, not a configuration change
Private cloudGovernment cloud programmes; residency with elasticityDepends on the provider's maturity, which varies enormously
Public cloudElastic demand, managed services, mature operations availableResidency and sovereignty; egress costs; skills; currency and budget volatility
HybridThe common real answer — central services in cloud, facility systems localTwo operating models to sustain
EdgeFacilities that must function while disconnectedSync, conflict resolution, device management at scale

Data residency is usually the first filter, and it is a legal question, not a technical one. Establish what the law and the health ministry's policy actually require — often "in-country" for identifiable data, with more latitude for aggregates — before evaluating providers. See GDPR where it applies, and national data protection law everywhere.

A frequent mistake: choosing a global public cloud region in a neighbouring country because it is nearest, without checking whether that is lawful. It usually is not, for identifiable clinical data.


Availability as a clinical parameter​

Availability targets should be set by clinicians against consequences, not chosen from a list.

SystemRealistic targetBecause
Emergency department EMRVery high; minutes of downtime matterCare stops
Facility EMRHigh during operating hoursCare degrades to paper
Interoperability layerHigh — everything routes through itBlocks all exchange
Client registryHigh — on the registration critical pathBlocks registration
Shared health recordMedium; degraded operation is tolerableClinicians work from local records
HMIS / reportingLower; hours of downtime are acceptableReports are late
Analytics / warehouseLowestNobody is harmed

Two consequences for design. First, tiering: not everything needs the same investment, and pretending otherwise means nothing gets it. Second, degraded mode — what the facility does when the central service is unreachable — is an architecture requirement for every high-tier system. Local caching of registries and store-and-forward queues at the edge are the standard answers. See offline-first.


Containers and orchestration​

Docker has become the default packaging format for health platforms — OpenMRS, DHIS2, HAPI FHIR, OpenHIM and most others publish images. That alone is a large improvement over hand-built servers.

Kubernetes is a separate decision, and a heavier one:

Adopt it when you run many services with differing scaling needs, have multiple teams deploying independently, need self-healing and rolling upgrades, and — critically — have people who can debug it.

Do not adopt it when you run three services on two servers with one system administrator. Docker Compose on well-managed virtual machines is a legitimate production architecture for a great many national deployments, and it is recoverable at 3 a.m. by someone who did not build it.

The failure mode is a ministry that adopts Kubernetes because it is best practice, and then cannot diagnose a failing pod during an outage. Operational capacity is a design input, not an aspiration.

Managed Kubernetes removes the control-plane burden but not the application operations burden, which is where most of the difficulty actually is.


Compute and storage patterns​

NeedOptionsHealth notes
Application runtimeVMs, containers, serverlessServerless suits event handlers and bulk processing; poor fit for stateful clinical platforms
Relational storagePostgreSQL (usual), MySQLMost health platforms assume one specifically; check before choosing
Object storageS3-compatibleDICOM images, bulk exports, backups
SearchElasticsearch / OpenSearchPatient search, terminology search, log analytics
CachingRedis / ValkeyTerminology expansions, session state
MessagingKafka, RabbitMQ, NATSEvent-driven exchange; see patterns

Storage growth is dominated by imaging. A radiology programme changes the infrastructure cost profile by an order of magnitude, and retention periods are measured in decades — tiered storage and a lifecycle policy are requirements, not optimisations.


Networking​

  • Segmentation. Clinical systems, medical devices, administrative networks and guest wifi must be separate. Medical devices are frequently unpatchable; isolation is the only available control.
  • Facility connectivity. Design for what exists — intermittent, low bandwidth, high latency — rather than for what is promised.
  • Private connectivity between the national data centre and cloud, where hybrid.
  • Egress control. Preventing data leaving is as important as preventing entry, and is more often overlooked.
  • DNS and certificates. Both are recurring outage causes. Automate renewal and monitor expiry; see security architecture.

Cloud providers​

Documented here for completeness. This knowledge base is not vendor-specific, and no provider is recommended over another.

  • AWS — HealthLake, HealthImaging, and general services
  • Azure — Azure Health Data Services, Microsoft Cloud for Healthcare
  • GCP — Cloud Healthcare API
  • Oracle Cloud, DigitalOcean, Hetzner, OVH and regional providers — often relevant where residency, cost or existing government agreements dominate

The managed FHIR/DICOM services are genuinely useful and create genuine lock-in. Evaluate the exit path — can you export everything in a standard format, and what would running the equivalent yourself cost? — as part of the selection, not afterwards.


In this section​


A sizing checklist​

  • Data residency requirements established, in writing, from legal counsel
  • Availability tier assigned per system, agreed with clinical leadership
  • Degraded mode designed for every high-tier system
  • RTO and RPO per system, set by consequence
  • Backup strategy including one immutable copy, with tested restores
  • Growth projection including imaging and audit logs
  • Operational team identified and staffed for the chosen complexity
  • Monitoring and alerting in place before go-live, not after
  • Disaster recovery site or region, and a rehearsed failover
  • Ten-year cost model including staff, not only hosting

References​